Papers with safety detection

3 papers
Reasoning’s Razor: Reasoning Improves Accuracy but Hurts Recall at Critical Operating Points in Safety and Hallucination Detection (2026.eacl-long)

Copied to clipboard

Challenge: a new study examines the suitability of reasoning for precision-sensitive classification tasks . false positives carry severe operational consequences, such as blocking legitimate queries .
Approach: They propose to use reasoning for classification tasks under low false positive rate regimes . they find that Think On improves overall accuracy, but performs poorly at low FPRs a .
Outcome: The proposed reasoning-augmented generation model outperforms self-verbalized confidence in precision-sensitive deployments.
InstructSafety: A Unified Framework for Building Multidimensional and Explainable Safety Detector through Instruction Tuning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing safety detection systems have limitations in terms of their versatility and interpretability.
Approach: They introduce a safety detection framework that unifies 7 common sub-tasks into a uniform formulation and process 39 human-annotated datasets for instruction tuning.
Outcome: The proposed framework unifies 7 common sub-tasks into a uniform formulation and then runs on 39 human-annotated datasets to fine-tune it.
False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has leveraged probing-based approaches to study the separability of malicious and benign inputs in Large Language Models’ internal representations.
Approach: They propose to use probing-based methods to study separability of malicious and benign inputs in LLMs' internal representations to detect harmful and benign content.
Outcome: The proposed methods show that they learn superficial patterns rather than semantic harmfulness.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations